Papers with standalone performance metrics

1 papers
To Err Is Human, How about Medical Large Language Models? Comparing Pre-trained Language Models for Medical Assessment Errors and Reliability (2024.lrec-main)

Copied to clipboard

Challenge: a 1999 report found that at least forty thousand deaths are a result of preventable medical errors.
Approach: They test pre-trained language models to characterize their error generation and reliability in medical assessment ability.
Outcome: The results show that pre-trained models can generate errors and perform better than human models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations